Papers with Text-to-SQL task

8 papers
BookSQL: A Large Scale Text-to-SQL Dataset for Accounting Domain (2024.naacl-long)

Copied to clipboard

Challenge: Existing models for accounting databases that can be queried using natural language are lacking in some domains.
Approach: They propose a large-scale text-to-SQL dataset for accounting and financial domains . they propose 'bookSQl' to be used to query accounting databases using natural language .
Outcome: The proposed model performs poorly on the existing model, pointing towards a more focused model for this domain.
SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for self-play fine-tuning do not generate new information and the large number of correct SQL queries produced by the opponent model reduces the main model’s ability to generate accurate SQL queries.
Approach: They propose a self-play fine-tuning method tailored for the Text-to-SQL task that synthesizes high-quality fine- tuning data iteratively based on the database schema and validation feedback to enhance model performance.
Outcome: The proposed method outperforms existing state-of-the-art methods on six open-source LLMs and five widely used benchmarks.
SDE-SQL: Enhancing Text-to-SQL Generation in Large Language Models via Self-Driven Exploration with SQL Probes (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches depend on static, pre-processed database information, which restricts the model’s capacity to deeply comprehend the underlying database content.
Approach: They propose a framework that empowers LLMs to perform Self-Driven Exploration of databases during inference.
Outcome: Evaluated on the BIRD benchmark with Qwen2.5-72B-Instruct, SDE-SQL achieves an 8.02 % improvement in execution accuracy over the baseline.
Enhancing Text-to-SQL Parsing through Question Rewriting and Execution-Guided Refinement (2024.findings-acl)

Copied to clipboard

Challenge: Existing prompt engineering methods exploit database content and execution feedback to improve text-to-sql performance.
Approach: They propose a framework for large language model-based text-to-sql task that exploits database content and execution feedback to improve execution accuracy.
Outcome: The proposed framework improves execution accuracy and usability by 12.41% and 5.38% on four widely used benchmarks.
Addressing Limitations of Encoder-Decoder Based Approach to Text-to-SQL (2022.coling-1)

Copied to clipboard

Challenge: Existing attempts on Text-to-SQL task show a dramatic decline in performance for new databases.
Approach: They propose a hybrid system that integrates rule-based and deep learning components to improve model accuracy.
Outcome: The proposed system achieves double-digit percentage improvement for non-Spider databases.
LitE-SQL: A Lightweight and Efficient Text-to-SQL Framework with Vector-based Schema Linking and Execution-Guided Self-Correction (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods rely on proprietary models to generate SQL queries.
Approach: They propose a lightweight framework that translates natural language questions into SQL queries.
Outcome: The proposed framework achieves 72.10% execution accuracy on BIRD and 88.45% on Spider 1.0 . it offers a practical solution for privacy-sensitive and resource-constrained settings.
Enhancing Text-to-SQL Capabilities of Large Language Models through Tailored Promptings (2024.lrec-main)

Copied to clipboard

Challenge: Large language models with prompting have achieved encouraging results on many natural language processing tasks due to the absence of task-tailored promptings.
Approach: They propose three promptings specifically designed for Text-to-SQL: SL-prompt, CC-promped, and SL+CC prompt.
Outcome: The proposed promptings achieve execution accuracy of 86.2% and test-suite accuracy of 76% . the granularity of schema linking and the order of clause generation have great impact on performance, which are considered little in previous research.
Enhancing Text-to-SQL Capabilities of Large Language Models: A Study on Prompt Design Strategies (2023.findings-emnlp)

Copied to clipboard

Challenge: In-context learning (ICL) is a new approach to natural language processing tasks that rely on large language models to make predictions based on context . recent studies have shown that neural symbolic design is the preferred choice for question answering systems because of its limited working memory and unreliable long-term memory.
Approach: They propose to extend in-context learning to question answering tasks that utilize structured knowledge sources and to explore various prompt design strategies for employing LLMs.
Outcome: The proposed approach outperforms the state-of-the-art system by 2.5 points and the best fine-tuned system by 5.1 points on the Spider dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations